We are seeking a Site Reliability Engineer to help build, maintain, and improve highly available infrastructure supporting critical business applications and services. This role sits within a collaborative engineering environment where reliability, automation, scalability, and performance are key priorities.
Responsibilities
• Design, build, and maintain highly available and scalable infrastructure
• Automate operational processes to improve efficiency and reduce manual intervention
• Manage and support Kubernetes-based containerized environments
• Develop and maintain Infrastructure as Code using Terraform and related tools
• Build and enhance CI/CD pipelines to streamline deployment processes
• Monitor system health, performance, and reliability across production environments
• Lead incident response efforts and drive root cause analysis for production issues
• Partner with engineering teams to improve system design, resilience, and observability
• Implement best practices around monitoring, alerting, capacity planning, and disaster recovery
• Continuously identify opportunities to improve reliability, performance, and operational excellence
Requirements
• Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent experience)
• Experience in Site Reliability Engineering, Platform Engineering, Production Engineering, DevOps, or Infrastructure Engineering
• Strong Linux systems administration experience
• Hands-on experience with Kubernetes and containerized environments
• Experience with Terraform or other Infrastructure as Code tools
• Strong knowledge of cloud platforms such as AWS, Azure, or GCP
• Proficiency in Python, Go, Bash, or similar scripting/programming languages
• Experience building and supporting CI/CD pipelines
• Familiarity with monitoring and observability tools such as Prometheus, Grafana, Datadog, Splunk, or ELK
• Strong troubleshooting and problem-solving skills in large-scale production environments